Papers with debiasing methods
Copied to clipboard
| Challenge: | Several methods have been proposed to mitigate bias in training on biased datasets. |
| Approach: | They propose to examine the effect of target class imbalance and stereotyping on model performance by analyzing binary classification, profession prediction and regression tasks. |
| Outcome: | The proposed methods show that data conditions have a strong influence on relative model performance. |
Copied to clipboard
| Challenge: | Existing approaches to assess and improve model fairness have been inconsistent and inconsistent. |
| Approach: | They propose an open-source python library for assessing and improving model fairness. |
| Outcome: | The proposed framework can be used for natural language, images, and audio. |
Copied to clipboard
| Challenge: | Recent advances in large language models have enabled impressive zero-shot capabilities across various natural language tasks. |
| Approach: | They propose two ways to exploit the emergent abilities of large language models for NLG assessment. |
| Outcome: | The proposed methods improve performance and positional biases in comparisons between candidates. |
Copied to clipboard
| Challenge: | Recent research has shown that distributional word vector spaces often encode stereotypical human biases, such as racism and sexism. |
| Approach: | They propose a platform that measures and mitigates bias in word embeddings by executing two (mutually composable) debiasing models. |
| Outcome: | The proposed platform can measure and mitiga bias in word embeddings. |
Copied to clipboard
| Challenge: | Existing benchmarks and resources for evaluating gender biases in multilingual settings are limited. |
| Approach: | They propose to extend DisCo to different Indian languages using human annotations to evaluate gender biases in multilingual models. |
| Outcome: | The proposed benchmarks and mitigation techniques are extended beyond English to evaluate gender biases in multilingual models. |
Copied to clipboard
| Challenge: | Pretrained Language Models (PLMs) are widely used in NLP for various tasks. |
| Approach: | They propose to modularly debias a pre-trained language model across multiple bias dimensions using structured knowledge and a large generative model. |
| Outcome: | The proposed model is able to debias a pre-trained language model across multiple bias dimensions in a semi-automated way. |
Copied to clipboard
| Challenge: | Existing methods for debiasing may generate incorrect or nonsensical predictions but leave aside individual commonsense facts, resulting in modified knowledge that elicits unreasonable or undesired predictions. |
| Approach: | They propose a framework that identifies encoding locations of biases within language models and then applies the Fairness-Stamp (FAST) they also propose 'BiaScope' to evaluate the retention of commonsense knowledge and generalization across paraphrased social biase. |
| Outcome: | The proposed framework surpasses state-of-the-art baselines with superior debiasing performance while not compromising the overall model capability for knowledge retention and prediction. |
Copied to clipboard
| Challenge: | Large-scale, pretrained vision-language models are growing in popularity due to impressive performance on downstream tasks with minimal finetuning. |
| Approach: | They propose to apply ranking metrics to image-text representations to investigate bias measures and debiasing methods to reduce various bias measures. |
| Outcome: | The proposed model reduces bias measures with minimal degradation to image-text representations. |
Copied to clipboard
| Challenge: | Existing methods to remove gender bias from word embeddings are insufficient, we argue . existing methods for gender-neutral modeling are ineffective, we conclude . |
| Approach: | They propose methods to reduce gender bias in word embeddings by debiasing them using text corpora. |
| Outcome: | The proposed methods show that they can reduce gender bias in word embeddings . the proposed methods are insufficient and should not be trusted, the authors argue . |
Copied to clipboard
| Challenge: | Recent research shows that dialogue systems trained on human conversation data are biased and can produce responses that reflect people’s gender prejudice. |
| Approach: | They propose a novel adversarial learning framework Debiased-Chat to train dialogue models free from gender bias while keeping their performance. |
| Outcome: | The proposed framework significantly reduces gender bias in dialogue models while maintaining the response quality. |
Copied to clipboard
| Challenge: | Existing approaches for debiasing datasets are weaker than current approaches for generalization. |
| Approach: | They propose a framework for analyzing multiple biases in training data to reduce bias weighting. |
| Outcome: | The proposed framework improves generalization on in-domain and out-of-domain datasets by weighting examples based on their strengths and bias strengths. |
Copied to clipboard
| Challenge: | racial descriptors alter embedding similarity scores and retrieval rankings, a new study shows . rife-specific biases can displace relevant records outside top-10 results, the study concludes . |
| Approach: | They propose to detect, measure, and mitigate racial bias in NLP systems deployed in criminal justice contexts . they propose to develop and evaluate debiasing techniques, validate synthetic findings on authentic law enforcement data . |
| Outcome: | The proposed research examines how bias propagates across retrieval pipelines . it shows that racial descriptors alter embedding similarity scores and retrieval rankings . |
Copied to clipboard
| Challenge: | a study of contextualised word embeddings shows discriminative biases are encoded in contextualised embeddables. |
| Approach: | They propose a fine-tuning method that can be applied at token- or sentence-levels to debias pre-trained contextualised embeddings. |
| Outcome: | The proposed method can be applied at token- or sentence-levels to debias pre-trained models without requiring retrains. |
Copied to clipboard
| Challenge: | Recent studies have shown that LLMs exhibit social biases inherited from training data. |
| Approach: | They propose a framework for evaluation and mitigation of bias in Large Language Models applied to complex clinical cases using a dataset based on the JAMA Clinical Challenge. |
| Outcome: | The proposed framework employs multiple choice questions and explanations to evaluate gender and ethnicity biases in LLMs. |
Copied to clipboard
| Challenge: | Recent debiasing methods in natural language understanding improve performance on out-of-distribution datasets by pressuring models into making unbiased predictions. |
| Approach: | They propose a general probing-based framework that allows for post-hoc interpretation of biases in language models and use an information-theoretic approach to measure the extractability of certain biase . |
| Outcome: | The proposed framework allows for post-hoc interpretation of biases in language models and measures the extractability of certain biase . |
Copied to clipboard
| Challenge: | Prior work has proposed debiasing methods that require human labelled examples, data augmentation and fine-tuning of LLMs, which are computationally expensive. |
| Approach: | They propose to suppress gender biases by providing textual preambles from manually designed templates and real-world statistics without accessing model parameters. |
| Outcome: | The proposed methods suppress gender biases in English LLMs using a CrowsPairs dataset without accessing model parameters. |
Copied to clipboard
| Challenge: | Existing work shows that Large Language Models (LLMs) are not robust to complex language understanding tasks due to reliance on spurious correlations of training datasets. |
| Approach: | They propose a method for measuring model reliance on spurious features by exploiting chosen biases on out-of-distribution (OOD) datasets. |
| Outcome: | The proposed method shows that the reported OOD gains of debiasing methods can't be explained by mitigated reliance on biased features, suggesting that biases are shared among different QA datasets. |
Copied to clipboard
| Challenge: | Existing approaches to debiase datasets rely on knowledge of bias attributes . current approaches focus on how to leverage kinds of supervision effectively . |
| Approach: | They propose to extend the supervision on bias by extending it into feature space. |
| Outcome: | Empirical results show that a low-dimensional subspace with intended features can represent biased datasets. |
Copied to clipboard
| Challenge: | Recent work has focused on measuring and mitigating bias in pretrained language models. |
| Approach: | They propose a dataset that measures and mitigates bias across gender,race, religion, and queerness . they compare REDDITBIAS to a widely used conversational DialoGPT model . |
| Outcome: | The proposed framework measures and mitigates bias across gender,race, religion, and queerness dimensions. |
Copied to clipboard
| Challenge: | Existing approaches to debiase Natural Language Understanding models use dataset biases instead of learning the intended task. |
| Approach: | They propose a debiasing framework that detects and purifies dataset biases using information entropy. |
| Outcome: | The proposed framework improves the stability of performance on out-of-distribution datasets for a set of widely adopted NLU models. |
Copied to clipboard
| Challenge: | Existing methods for debiasing large language models require external bias knowledge or annotated non-biased samples, which is lacking for position debiases. |
| Approach: | They propose a self-supervised position debiasing framework that leverages unsupervised responses from pre-trained LLMs for debiazing without external bias knowledge. |
| Outcome: | The proposed framework outperforms existing methods in mitigating three types of position biases on eight datasets and five tasks. |
Copied to clipboard
| Challenge: | Existing methods to debiase samples with biased features obstructs the model in learning from non-biased parts of the samples. |
| Approach: | They propose to eliminate spurious correlations in a fine-grained manner from a feature space perspective by using Random Fourier Features and weighted re-sampling to decorrelate dependencies between features. |
| Outcome: | The proposed method eliminates spurious correlations in a fine-grained manner from a feature space perspective. |
Copied to clipboard
| Challenge: | a geo-cultural gap in NLP evaluation hinders evaluation of societal biases . authors propose a new method to collect stereotypes from large language models . |
| Approach: | They propose a new method that integrates sourcing and validation of existing data into a single workflow. |
| Outcome: | The proposed method improves LACES by integrating new stereotype entries and validation of existing data. |
Copied to clipboard
| Challenge: | Existing methods to develop meta-embeddings from source embeddings contain unfair gender-related biases, and how these influence the meta-bedding has not been studied yet. |
| Approach: | They propose to use multiple debiasing methods on a single source embedding to create a gender-based meta-embedding. |
| Outcome: | The proposed method amplifies gender biases compared to input source embeddings. |
Copied to clipboard
| Challenge: | Existing methods for debiasing multimodal models use approximate heuristics to represent the biases, such as shallow features from early stages of training or unimodal features for multimodal tasks like VQA, which may not be accurate. |
| Approach: | They propose a method that leverages causally-motivated information minimization to learn the confounder representations of a causal graph for multimodal data. |
| Outcome: | The proposed method improves out-of-distribution performance on multiple multimodal datasets without sacrificing in-distance performance. |
Copied to clipboard
| Challenge: | Existing methods for debiasing toxic language data are limited in their ability to prevent biased behavior in toxic language detection systems. |
| Approach: | They propose to debiase toxic language detection models using lexical and dialectal markers using synthetic labels instead of traditional methods. |
| Outcome: | The proposed method reduces dialectal associations with toxicity despite the use of synthetic labels . |
Copied to clipboard
| Challenge: | Pretrained language models (PLMs) propagate social stigmas and stereotypes, a critical concern given their widespread use. |
| Approach: | They adapt two intrinsic bias benchmarks to quantify racial and LGBTQ+ biases in prevalent PLMs and empirically evaluate the effectiveness of various debiasing methods in mitigating these biase. |
| Outcome: | The proposed methods reduce biases without compromising performance in downstream tasks. |
Copied to clipboard
| Challenge: | Recent research shows word embeddings have strong gender biases in embeddable spaces . a proposed method can be used to debiase word embeds without loss of semantic information . |
| Approach: | They propose a latent disentanglement method with a siamese auto-encoder structure with an adapted gradient reversal layer to debiase word embeddings. |
| Outcome: | The proposed method can preserve semantic information during debiasing while minimizing loss of semantic information for extrinsic NLP tasks. |
Copied to clipboard
| Challenge: | Existing methods to reduce bias have been shown to be effective over real-world datasets. |
| Approach: | They propose two new training objectives which directly optimise for the widely-used criterion of equal opportunity. |
| Outcome: | The proposed training objectives directly optimise for the widely-used criterion of equal opportunity while maintaining high performance over two classification tasks. |
Copied to clipboard
| Challenge: | Recent studies have shown that many well-developed Visual Question Answering systems suffer from bias problem. |
| Approach: | They propose a way to mitigate bias problem by subtracting bias score from standard VQA base score. |
| Outcome: | The proposed method improves on the VQA v2.0 and VQA-CP V2,0 datasets. |
Copied to clipboard
| Challenge: | Recent studies have shown that AI is unfair in many real-world applications such as computer vision and recommendations. |
| Approach: | They propose to use a benchmark dataset to study the fairness of dialogue systems to understand their bias. |
| Outcome: | The proposed methods reduce the bias in dialogue systems significantly. |
Copied to clipboard
| Challenge: | In recent years, NLP methods have found increasing adoption in the social sciences . however, CSS must be crucially interested in the algorithmic fairness of the underlying methods . |
| Approach: | They propose two methods which mask proper names and pronouns during training of the model, thus removing personal information bias. |
| Outcome: | The proposed methods decrease frequency bias while keeping the overall performance stable. |
Copied to clipboard
| Challenge: | Existing studies utilize social media platforms such as Twitter to build models for crisis event analysis, but semi-supervised approaches require annotating vast amounts of data and are impractical due to limited response time. |
| Approach: | They propose a method that stores and performs equal sampling for generated pseudo-labels from each class at each training iteration. |
| Outcome: | The proposed method performs better than existing methods in both in-distribution and out-of-difference settings. |
Copied to clipboard
| Challenge: | Existing debiasing methods modify all of the PLM parameters, which is costly and leads to (catastrophic) forgetting of useful language knowledge. |
| Approach: | They propose a modular debiasing approach based on dedicated adapters that inject adapter modules into the original PLM layers and update only the adapters. |
| Outcome: | The proposed approach is based on dedicated adapters and retains fairness even after large-scale training. |
Copied to clipboard
| Challenge: | Parameter-efficient fine-tuning (PEFT) addresses the memory footprint issue of full fine- tuning by modifying only a subset of model parameters. |
| Approach: | They propose a framework that debiases models in a biased-to-unbiased order and uses only a subset of parameters to modify model parameters. |
| Outcome: | The proposed framework accelerates convergence on unbiased examples by approximately twofold and improves ID and OOD performance by 1.2% and 8.0%, respectively. |
Copied to clipboard
| Challenge: | Existing studies on bias dataset construction and mitigation focus on one demographic group . in real-world applications, there are more than two demographic groups at risk of the same bias. |
| Approach: | They propose to analyze and reduce biases across multiple demographic groups using a multi-demographic bias dataset. |
| Outcome: | The proposed method can mitigate biases among multiple demographic groups effectively, the authors show . |
Copied to clipboard
| Challenge: | Existing methods for stance detection are task-agnostic, which fail to utilize task knowledge to better discriminate between genuine and bias features. |
| Approach: | They propose to incorporate stance reasoning process as task knowledge to aid in learning genuine features without using targets. |
| Outcome: | The proposed model achieves better performance than previous task-agnostic debiasing methods on new test sets. |
Copied to clipboard
| Challenge: | Existing methods for debiasing depend on attribute labels and target attributes. |
| Approach: | They propose a method that uses class-wise variance of embeddings to reduce the effects of debiasing on a downstream task. |
| Outcome: | The proposed method outperforms baselines that rely on attribute labels while maintaining performance on the target task. |
Copied to clipboard
| Challenge: | Recent proposed debiasing methods rely on the assumption that the types of bias should be known a-priori, which limits their application to many NLU tasks and datasets. |
| Approach: | They propose a framework that prevents models from mainly utilizing biases without knowing them in advance. |
| Outcome: | The proposed framework allows existing methods to retain performance improvement on challenge datasets without specifically targeting biases. |
Copied to clipboard
| Challenge: | 'spurious correlations' have been used in NLP to informally denote any undesirable feature-label correlations. |
| Approach: | They formalize this distinction using a causal model and probabilities of necessity and sufficiency, which delineates causal relations between a feature and a label. |
| Outcome: | The proposed model is invariant to the feature, but not sufficient for prediction. |
Copied to clipboard
| Challenge: | Existing approaches to mitigate the detrimental effect of bias on the network include debiasing methods that down-weight the biased examples identified by an auxiliary model, which is trained with explicit bias labels. |
| Approach: | They propose a framework that introduces binary classifiers between the auxiliary model and main model, coined bias experts, to reduce the detrimental effect of bias on the network. |
| Outcome: | The proposed approach outperforms the state-of-the-art on various datasets while achieving high performance on in-distribution data. |
Copied to clipboard
| Challenge: | Recent advances in Multi-modal Language Models have shown remarkable performance in multimodal tasks . however, these models often exhibit inherent biases that compromise their reliability and fairness. |
| Approach: | They propose a framework that integrates Distill, Dynamic Drop, and Merge to address these challenges. |
| Outcome: | The proposed framework outperforms existing methods in balancing debiasing and improving performance on the MMSD2.0 sarcasm detection dataset. |
Copied to clipboard
| Challenge: | Existing methods for debiasing are ineffective in addressing the reverse word-overlap bias. |
| Approach: | They propose to investigate the reverse word-overlap bias in NLI models . they find that existing debiasing methods are generally ineffective . |
| Outcome: | The proposed model is biased towards the non-entailment label on instances with low overlap . the proposed model does not have minority examples, the authors show . |
Copied to clipboard
| Challenge: | Existing studies on social biases in language models have focused on only English. |
| Approach: | They propose to use a Chinese dataset for bias evaluation and mitigation of Chinese conversational language models. |
| Outcome: | The proposed dataset includes under-explored bias categories, such as ageism and appearance biases, which received less attention in previous studies. |
Copied to clipboard
| Challenge: | Recent studies have shown that strong natural language understanding models are prone to relying on unwanted dataset biases without learning the underlying task. |
| Approach: | They propose two learning strategies to train neural models that are more robust to dataset biases and transfer better to out-of-domain datasets. |
| Outcome: | The proposed methods improve robustness in all settings and transfer better to out-of-domain datasets. |
Copied to clipboard
| Challenge: | Recent studies show that pre-trained language models rely heavily on idiosyncratic biases of datasets. |
| Approach: | They propose a method which discourages models from exploiting biases while enabling them to receive enough incentive to learn from all the training examples. |
| Outcome: | The proposed method improves on out-of-distribution datasets while maintaining original in-district accuracy. |
Copied to clipboard
| Challenge: | Empirical evaluations of large language models demonstrate that they improve performance in a wide range of tasks. |
| Approach: | They propose a label-free method for mitigating selection bias during inference by reformulating debiasing as an optimization task. |
| Outcome: | The proposed method mitigates selection bias and improves performance compared to existing methods. |
Copied to clipboard
| Challenge: | Existing methods for debiasing large language models incur high human and computational costs and are limited in their effectiveness. |
| Approach: | They propose a model-agnostic, inference-time debiasing framework that enforces fairness by filtering generation outputs in real time. |
| Outcome: | The proposed framework mitigates social bias across a range of LLMs while preserving overall generation quality. |
Copied to clipboard
| Challenge: | Recent studies have demonstrated that large language models exhibit social biases . however, debiasing methods may degrade the capabilities of LLMs if they are not properly evaluated . |
| Approach: | They propose a Japanese benchmark to evaluate social biases and cultural commonsense in large language models in a unified format. |
| Outcome: | The proposed method degrades the performance of the LLMs on the cultural commonsense task by 75%. |
Copied to clipboard
| Challenge: | Large language models (LLMs) are expensive yet powerful ways to annotate text, and can be inconsistent when compared with experts. |
| Approach: | They propose to combine LLM annotations with a limited number of expensive expert annotations to produce valid estimates. |
| Outcome: | The proposed methods produce consistent estimates under theoretical assumptions, but they are not comparable across finite datasets. |
Copied to clipboard
| Challenge: | Existing debiasing frameworks can detect known dataset biases and spurious correlations in data. |
| Approach: | They propose a framework that learns to be undecided in its predictions for data samples . they propose 'contrary' objective that learn debiased and robust representations from biased views . |
| Outcome: | The proposed framework outperforms existing methods against out-of-domain and hard test samples without compromising performance. |
Copied to clipboard
| Challenge: | Existing methods for debiasing use uniform bias corrections across all input queries . weak debiases retains bias in sensitive queries, while weak dealiases in biased ones . |
| Approach: | They propose a framework that selectively applies debiasing based on input sensitivity . RG-TTA adaptively triggers fairness regularization based upon bias sensitivity of each input . |
| Outcome: | Experiments show that debiasing improves zero-shot performance while maintaining fairness . weak debiased queries distort semantically meaningful information while weak ones fail to mitigate stereotypes . |
Copied to clipboard
| Challenge: | Existing debiasing methods improve overall fairness, but fail to reduce framing-induced disparities. |
| Approach: | They propose a framing-aware debiasing method that encourages LLMs to be more consistent across frams. |
| Outcome: | The proposed method reduces overall bias and improves robustness against framing disparities, enabling LLMs to produce fairer and more consistent responses. |
Copied to clipboard
| Challenge: | Existing debiasing methods create biased responses by completely removing an entire modality, forming an extreme and static training environment. |
| Approach: | They propose a method to debiase multimodal large language models by masking one modality and then enlarge the margin between clean and adversarial responses. |
| Outcome: | The proposed method achieves superior debiasing performance while maintaining general capabilities. |